iT邦幫忙

2026 iThome 鐵人賽

DAY 14
0
自我挑戰組

Data Engineer 下班後偷學 AI系列 第 14 篇

GraphRAG in Milvus

  • 分享至 

  • xImage
  •  

前幾天介紹的 RAG,大多還是圍繞在「找到和 Query 最相似的 Chunk」,但有些問題的答案並不會直接存在某個 Chunk 裡,而是分散在不同文件,必須先理解資料之間的關係才能找到答案。這種需要跨越多個 Entity 與 Relation 才能得到答案的問題,通常稱為 Multi-hop Question,也正是 GraphRAG 想處理的場景之一。今天的程式內容與介紹是參考 Milvus 的官方 Graph RAG 範例。

GraphRAG

一般 RAG 主要依靠 Vector Similarity:

Query
→ Embedding
→ Vector Search
→ Relevant Chunks
→ LLM

GraphRAG 則會先從文件中整理出 Entity 和 Relation,建立一個 Knowledge Graph。

例如:

Euler ──student_of──> Johann Bernoulli

Daniel Bernoulli ──son_of──> Johann Bernoulli

Daniel Bernoulli ──contributed_to──> Fluid Dynamics

一條 Relation 通常可以表示成:

Subject → Predicate → Object

也就是常見的 Triple:

(Daniel Bernoulli, contributed_to, Fluid Dynamics)

所以 GraphRAG 的 Retrieval 不再只有「這段文字跟 Query 像不像」,而是可以先找到 Query 相關的 Entity,再沿著 Relation 向外展開,取得原本在 Vector Space 中不一定很接近、但在 Knowledge Graph 上有關聯的資料。

簡化後的流程大概是:

Document
→ Entity / Relation Extraction
→ Knowledge Graph
→ Retrieval
→ Graph Expansion
→ Rerank
→ Relevant Passages
→ LLM

GraphRAG 並不是要取代 Vector Search,而是在原本的 Retrieval 上加入 Graph Relationship。

GraphRAG with Milvus

Milvus 本身並不是 Graph Database,沒有 Neo4j 那種原生 Graph Traversal 或 Cypher Query。

Milvus 官方的做法,是把 GraphRAG 拆成:

  • Entity
  • Relation
  • Passage

並分別建立 Collection。

Entity 和 Relation 都可以轉成 Embedding 存進 Milvus,Graph 本身的關聯則另外用 Entity → Relation、Relation → Passage 的 mapping 保存。官方範例會同時對 Entity 和 Relation 做 Vector Search,再根據搜尋結果向外展開 Subgraph。

例如文件中有:

dataset = [
    {
        "passage": "Euler was a student of Johann Bernoulli.",
        "triplets": [
            ["Euler", "was a student of", "Johann Bernoulli"]
        ]
    },
    {
        "passage": "Daniel Bernoulli was the son of Johann Bernoulli.",
        "triplets": [
            ["Daniel Bernoulli", "was the son of", "Johann Bernoulli"]
        ]
    },
    {
        "passage": "Daniel Bernoulli made major contributions to fluid dynamics.",
        "triplets": [
            ["Daniel Bernoulli", "made major contributions to", "fluid dynamics"]
        ]
    }
]

實際系統通常會使用 LLM、NER 或 Information Extraction Model 從文件中抽出 Triplet,這邊先假設已經處理完成。

接著把 Entity、Relation、Passage 分開存進 Milvus:

milvus_client = MilvusClient(uri="./milvus.db")

entity_col_name = "entity_collection"
relation_col_name = "relation_collection"
passage_col_name = "passage_collection"

for collection_name in [
    entity_col_name,
    relation_col_name,
    passage_col_name,
]:
    milvus_client.create_collection(
        collection_name=collection_name,
        dimension=embedding_dim,
    )

./milvus.db 代表直接使用 Milvus Lite,所以這個 GraphRAG 範例不需要另外架 Milvus Server。官方目前的 Graph RAG 教學也是使用 Milvus Lite。

真正 Query 時,會有兩條 Retrieval Path。

第一條是從 Query 中先抽 Entity:

query = "What contribution did the son of Euler's teacher make?"

query_entities = ["Euler"]

entity_search_res = milvus_client.search(
    collection_name=entity_col_name,
    data=[
        embedding_model.embed_query(entity)
        for entity in query_entities
    ],
    limit=3,
    output_fields=["id"],
)

另一條則直接拿完整 Query 搜尋 Relation:

query_embedding = embedding_model.embed_query(query)

relation_search_res = milvus_client.search(
    collection_name=relation_col_name,
    data=[query_embedding],
    limit=3,
    output_fields=["id"],
)

也就是:

Query Entity
→ Entity Search
→ Related Entities

Query
→ Relation Search
→ Related Relations

Milvus 負責的主要還是我們前幾天一直在做的 Vector Retrieval。

Subgraph Expansion

接下來才是 GraphRAG 最重要的部分。

假設 Entity Search 找到了:

Euler

從 Graph 中可以知道:

Euler
→ Johann Bernoulli

再往下一層:

Johann Bernoulli
→ Daniel Bernoulli

最後:

Daniel Bernoulli
→ Fluid Dynamics

這個向外搜尋 Graph 的過程就是 Subgraph Expansion。

Milvus 官方範例使用 Entity-Relation Adjacency Matrix 來保存 Graph 關係:

entity_relation_adj = np.zeros(
    (len(entities), len(relations))
)

for entity_id in range(len(entities)):
    entity_relation_adj[
        entity_id,
        entityid_2_relationids[entity_id]
    ] = 1

如果某個 Entity 和某個 Relation 有關係,對應位置就是 1。

之後就可以透過 Matrix Multiplication 算出 Entity 之間的連接:

entity_adj = (
    entity_relation_adj
    @ entity_relation_adj.T
)

再往外展開一層,就是下一個 Hop。

官方範例也是透過 adjacency matrix 做指定 degree 的 Subgraph Expansion,取得 Entity Search 和 Relation Search 周圍的 Candidate Relations。

Reranking

Graph Expansion 還有一個很明顯的問題:

Graph 越往外展開,Relation 只會越來越多。

例如找到某個人物後,他可能同時具有:

father_of
worked_with
born_in
studied_at
wrote
invested_in
...

但這些 Relation 不一定都和 Query 有關。

所以 Subgraph Expansion 後,通常還需要一層 Reranking:

Entity / Relation Search
→ Subgraph Expansion
→ Candidate Relations
→ Reranking
→ Relevant Relations

Milvus 官方範例使用 LLM 對 Candidate Relations 再做一次篩選,最後把留下的 Relation 對應回原本的 Passage,再把 Passage 當成 Context 交給 LLM。

所以完整流程大概可以整理成:

Query
→ Entity / Relation Retrieval
→ Subgraph Expansion
→ Relation Reranking
→ Passage Retrieval
→ LLM

看到這裡其實會發現,GraphRAG 並不是一個完全不同的東西。

前幾天提到的:

Vector Search
Multi-way Retrieval
Reranking

全部還是存在,只是今天在 Retrieval 中間又加入了一層 Graph Expansion。

GraphRAG 的代價

GraphRAG 的優勢是可以處理一般 Vector Search 不擅長的 Multi-hop 問題,但代價也很明顯。

普通 RAG 可能只需要:

Chunk
→ Embedding
→ Search

GraphRAG 還需要處理:

Entity Extraction
Relation Extraction
Entity Resolution
Graph Construction
Graph Retrieval

尤其是 Entity Resolution。

例如:

OpenAI
Open AI
OpenAI Inc.

到底是不是同一個 Entity?

如果沒有處理好,Knowledge Graph 本身就是錯的,後面的 Vector Search、Graph Expansion 和 Reranking 做得再好也沒用。

所以 GraphRAG 比較適合資料本身就具有大量 Relationship,而且 Query 常常需要 Multi-hop Reasoning 的場景,例如人物關係、公司與供應鏈、論文引用、法律案件或其他 Knowledge Base。

如果只是一般 FAQ 或單一文件內容查詢,傳統 RAG 通常反而簡單很多。

每日一句

人生好難。


上一篇
Milvus 的 Filtered Search 以及背後的設計思考
下一篇
GraphRAG 補完計畫
系列文
Data Engineer 下班後偷學 AI 共 16 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言